Unified Real-Time Data Orchestration System with AI-Powered Predictive Compute Allocation for Cross-Domain Enterprise Operations

Authors: Surender Kusumba Trinamix Inc., USA

                         Phanindra Gangina Awoit Systems Inc, USA

Abstract

Correlated enterprise streams can propagate capacity pressure across order, payment, and customer-service domains, while reactive autoscaling responds only after queues form. This study specifies a unified orchestration architecture that combines event-time telemetry, cross-domain forecasting, uncertainty-aware planning, and constrained horizontal replica allocation. The evidence is obtained exclusively from a deterministic, discrete-time computational simulation; the evaluation did not start containers, execute Kafka or Flink processes, operate a Kubernetes cluster, or invoke a Kubernetes autoscaling API. The simulator compared conservative static allocation, tuned reactive control, and the predictive strategy across six synthetic workload families in 450 paired runs covering pilot, confirmatory, drift, ablation, and failure conditions. Across ramp, burst, and cascade workloads, predictive allocation reduced median paired p95 latency by 49.5% (95% CI [10.2%, 70.4%]) and SLA violations by 46.4% (95% CI [43.7%, 49.8%]) versus reactive control. It reduced normalized requested replica capacity by 29.8% versus static allocation while preserving simulated throughput. Cross-domain features were beneficial in cascades but not simultaneous bursts, and concept drift remained a limitation. These findings provide simulation-level evidence for predictive horizontal scaling; they do not establish production readiness or physical Kubernetes effectiveness.

Keywords: real-time data orchestration; predictive autoscaling; workload forecasting; cross-domain enterprise operations; computational simulation; Kubernetes; resource allocation

1. Introduction

Enterprise operations increasingly depend on continuous data flows rather than isolated batch transfers. An order may initiate inventory reservation, payment authorization, ledger posting, fulfillment updates, and customer inquiries within a short interval. Each domain has its own schema, latency expectation, processing path, and failure mode, yet the workloads are operationally connected. A promotion, billing cycle, or upstream incident can therefore cause a sequence of correlated load changes across several services. Stream-processing platforms provide event-time semantics, state management, durable replay, and low-latency transformations for such workloads, while container orchestrators provide placement, recovery, and horizontal elasticity (Akidau et al., 2015; Burns et al., 2016; Carbone et al., 2015). These capabilities are essential, but their control boundaries are often separated.

In a conventional deployment, a data pipeline observes input rate, lag, watermarks, backpressure, and processing latency, whereas a platform autoscaler observes CPU utilization or another threshold metric. The pipeline can identify that a queue is growing, but it does not necessarily determine how much compute capacity should be made ready. The autoscaler can add replicas, but it commonly acts only after utilization or backlog has crossed a configured threshold. Provisioning is not instantaneous: a scheduler must place a pod, pull or access its image, initialize the runtime, restore state when necessary, and wait for readiness. A reactive decision can therefore be technically correct and still arrive too late to prevent a temporary increase in tail latency. Surveys of cloud elasticity and autoscaling repeatedly identify decision signals, response delay, stability, and the trade-off between service quality and unused capacity as central concerns (Al-Dhuraibi et al., 2018; Lorido-Botran et al., 2014; Qu et al., 2018).

The problem becomes more difficult when related domains share a cluster. Scaling each domain independently ignores the possibility that one stream is a leading signal for another. For example, an order surge can precede payment traffic by tens of seconds and support traffic by several minutes. Domain-isolated control waits for each downstream workload to manifest its own pressure. A cross-domain controller can instead treat the upstream pattern as evidence about future demand, provided that event correlation is observable and the relationship is not assumed to be permanent. This distinction is important because accurate forecasting alone does not guarantee better system behavior. A forecast must be available before the actuation deadline, translated into feasible resources, bounded by cluster capacity, protected against uncertainty, and prevented from causing repeated scale reversals. Recent work on proactive microservice scaling, resource-efficient horizontal scaling, and stream-processing autoscalers strengthens the case for workload-aware control while also showing that dependency, state, and reconfiguration costs remain material (Ahmad et al., 2024; Hossen et al., 2022; Luo et al., 2022; Siachamis et al., 2024).

This study addresses the separation between data orchestration and compute control by specifying one auditable horizontal-scaling loop. The target implementation architecture connects a Kafka and Apache Flink data plane to a telemetry and feature plane, a model-agnostic forecast service, a constrained multi-objective allocator, a policy guard, and a Kubernetes controller interface. The computational artifact does not instantiate those platforms; it represents their decision-relevant states through deterministic workload, queue, readiness-delay, forecast, and control modules. At each 15-second cycle, the artifact validates signal freshness, predicts demand over 30-, 60-, and 120-second horizons, selects a bounded replica target, applies safety and stability constraints, and records the simulated outcome using immutable run and decision identifiers. CPU and memory requests per replica remain fixed, and no vertical-scaling or slow-loop action is evaluated. Stale telemetry, forecast failure, high uncertainty, capacity exhaustion, and authority conflict trigger a hold-safe plan or tuned reactive fallback.

The artifact is studied through Design Science Research (DSR) and a controlled deterministic computational simulation, not through a physical or virtual Kubernetes testbed. DSR links the operational problem to the architecture and its evaluation (Hevner et al., 2004; Peffers et al., 2007). No broker, stream-processing job, container scheduler, pod, Horizontal Pod Autoscaler, or Kubernetes API is executed. Instead, the discrete-time model represents three enterprise-like domains (order-inventory, payment-finance, and customer service) and compares conservative static allocation, tuned reactive control, and the proposed predictive horizontal strategy under stable, periodic, ramp, sudden-burst, cross-domain-cascade, and concept-drift workloads. Workload and service trajectories are paired by seed, all treatment state is reset between runs, and claims are governed by prespecified performance, efficiency, stability, throughput, and reproducibility gates that apply only within the frozen simulation.

The study addresses five research questions:

  1. RQ1: How can heterogeneous enterprise data streams and compute-control signals be integrated into a unified real-time orchestration architecture?
  2. RQ2: How accurately and efficiently can a domain-aware AI model predict short-term compute demand under stable, periodic, ramp, burst, cascade, and drift workloads?
  3. RQ3: Within the deterministic simulation, to what extent does predictive horizontal replica allocation improve tail latency, throughput, and service-level compliance relative to static allocation and tuned reactive autoscaling?
  4. RQ4: What requested-replica efficiency and scaling-stability trade-offs arise across simulated workload patterns?
  5. RQ5: Which simulated contributions arise from cross-domain signals, uncertainty handling, and constrained horizontal allocation?

The paper makes four contributions. First, it specifies a unified data-and-compute control plane whose decisions remain traceable from event observation to a horizontal replica target. Second, it defines a domain-aware forecast contract that includes uncertainty, freshness, inference overhead, and model identity rather than only a point estimate. Third, it introduces a constrained horizontal allocation policy that combines predicted service-level risk, requested-capacity cost, scaling churn, and domain balance under explicit safety constraints. Fourth, it provides a reproducible 450-run computational benchmark linking forecast quality, simulated service outcomes, requested-capacity efficiency, ablations, failure emulation, and joint claim gates. These contributions are bounded to the deterministic model; physical Kafka, Flink, container, and Kubernetes validation remains future work.

2. Related Work

2.1 Event-Time Stream Processing and Data Orchestration

Modern stream processors treat an unbounded stream as a continuously evolving dataset whose correctness depends on time, state, triggering, and late-data policy. The Dataflow model formalized the separation among event time, processing time, windowing, triggers, and accumulation, providing a useful basis for reasoning about out-of-order enterprise events (Akidau et al., 2015). Apache Flink operationalizes similar principles through stateful operators, event-time windows, watermarks, checkpoints, and recovery (Carbone et al., 2015). Kafka contributes a durable partitioned log that decouples producers and consumers and supports deterministic replay when offsets, partitions, and schemas are controlled (Kreps et al., 2011). These foundations explain how domain events can be processed continuously, but they do not by themselves determine the compute capacity that should be available at a future point.

Elastic stream-processing research has examined operator parallelism, state movement, latency constraints, and adaptation cost. Gedik et al. (2014) studied elastic scaling of data-flow regions, while Lohrmann et al. (2015) linked elastic resource decisions to latency guarantees. Dhalion introduced self-regulation for Heron and showed the value of a control architecture that diagnoses symptoms, proposes actions, and verifies results (Floratou et al., 2017). Operational work at LinkedIn further demonstrated that resource sizing for large portfolios of streaming applications is a recurring platform problem rather than a one-time deployment choice (Singh et al., 2020). A recent evaluation of stream-processing autoscalers also emphasizes that performance must be tested under diverse dynamic workloads and that a controller’s convergence and stability cannot be inferred from one steady-state scenario (Siachamis et al., 2024).

These studies motivate a feedback loop, but the present work adopts a broader boundary. Event and telemetry contracts from three business domains are treated as first-class inputs to compute decisions. The proposed system retains stream-processing semantics in the data plane while placing forecasting, allocation, and actuation in a separate control plane. This separation allows the benchmark to measure AI overhead and to keep event ingestion durable when predictive control is degraded.

2.2 Reactive, Proactive, and Learning-Based Autoscaling

Autoscaling methods can be organized by decision timing, signal type, and action space. Static provisioning selects capacity before a workload begins and is operationally simple, but it can waste resources or fail under demand outside the sizing envelope. Reactive threshold control changes capacity after a measured signal departs from a target. It remains widely used because its behavior is understandable and its input requirements are modest. Predictive control estimates future workload or utilization and acts before the expected pressure arrives. Hybrid approaches combine a prediction with reactive protection or fallback. Reviews show that no single category is universally superior because benefit depends on workload predictability, provisioning delay, model overhead, error tolerance, and policy stability (Lorido-Botran et al., 2014; Qu et al., 2018).

Learning-based resource management ranges from supervised forecasting to reinforcement learning. Resource Central showed that workload characterization and prediction can improve large-scale cloud resource decisions when production telemetry is made useful to schedulers (Cortez et al., 2017). Deep reinforcement learning has also been applied to resource management, illustrating that a controller can learn sequential allocation policies, although training cost, interpretability, transfer, and safety complicate direct operational use (Mao et al., 2016). Gradient-boosted trees provide a lower-complexity alternative for tabular lag, rolling, calendar, and cross-domain features (Chen & Guestrin, 2016), whereas recurrent models such as long short-term memory can represent temporal dependencies when sufficient ordered data exist (Hochreiter & Schmidhuber, 1997).

Recent microservice research moves from per-container thresholds toward workload and dependency awareness. PEMA searches for resource-efficient allocations while protecting quality of service and adapts without requiring a large offline training set (Hossen et al., 2022). Madu predicts individual microservice workload and explicitly represents uncertainty and burst risk, connecting forecast design to proactive scaling decisions (Luo et al., 2022). Smart HPA coordinates resource exchange under constrained capacity rather than treating each service as an isolated claimant (Ahmad et al., 2024). These systems demonstrate complementary mechanisms, but their evaluation objects differ. Forecast accuracy, resource reduction, service dependencies, and cluster coordination are not always studied together with event-time orchestration, correlated business-domain signals, and a complete event-to-decision audit trail.

2.3 Horizontal Scaling and Multi-Objective Control

Horizontal scaling changes the number of replicas, whereas vertical scaling changes CPU or memory requests and may require pod replacement. These actions have different operational costs and should not be treated as interchangeable. The present study restricts its action space to horizontal replicas because the simulator represents readiness delay but does not model container restart, pod replacement, state restoration, or rollout disruption. Within that boundary, the controller still faces competing objectives: additional replicas can reduce service-level risk but increase requested capacity, while aggressive scale-in can amplify prediction error, cause oscillation, or disadvantage a lower-volume domain.

Cloud-elasticity literature consequently frames autoscaling as a constrained control problem rather than a single-metric trigger (Al-Dhuraibi et al., 2018; Qu et al., 2018). Resource-management surveys for distributed stream processing similarly distinguish provisioning, placement, scheduling, and reconfiguration decisions and identify state, topology, and heterogeneity as recurring concerns (Liu & Buyya, 2020). The allocator evaluated here represents service-level risk, requested replica capacity, scaling churn, and domain imbalance as separate objective terms while enforcing hard capacity and safety constraints. CPU and memory requests are fixed coefficients used to account for requested capacity, not control variables. Vertical right-sizing is therefore excluded from the current evidence and reserved for future work.

2.4 Cross-Domain Signals and End-to-End Evidence

A multi-service system contains temporal and structural dependencies. Traffic can propagate through a call graph, an event graph, or an enterprise process even when services have different external interfaces. Dependency-aware prediction can exploit leading information, but correlations may weaken during concept drift or change after a business process is redesigned. Cross-domain signals must therefore be time-bounded, versioned, and accompanied by uncertainty. The controller should not assume that an upstream surge always produces the same downstream magnitude or delay.

The central evidence gap is not the absence of forecasting or autoscaling studies. It is the limited joint evaluation of five elements: unified event and control telemetry, cross-domain short-horizon prediction, constrained horizontal replica allocation, simulated end-to-end decision execution, and reproducible decision provenance. A model can attain low average error while missing a burst that determines the service outcome. A controller can reduce latency by overprovisioning. A resource-efficient policy can violate throughput. A favorable average can hide an unstable domain or repeated scale reversals. The present design evaluates this computational chain from forecast to outcome and requires performance, efficiency, stability, guardrail, and reproducibility evidence before any simulation-level superiority claim is permitted.

Table 1. Positioning of representative related work

WorkStream/event contextPredictive signalDependency or cross-domain scopeConstrained allocationExecuted end-to-end evaluationReproducibility emphasis
Akidau et al. (2015)YesNoNoNoData-processing focusConceptual/implementation detail
Floratou et al. (2017)YesDiagnostic controlTopology-awareYesYesOperational evidence
Hossen et al. (2022)MicroservicesAdaptive workload signalService interactionsQoS/resource constraintsYesPrototype evaluation
Luo et al. (2022)MicroservicesUncertainty-awarePer-service learningProactive scalingYesExperimental
Ahmad et al. (2024)MicroservicesReactive/heuristicCoordinated servicesCapacity-awareYesCode available
Siachamis et al. (2024)Stream processingController-dependentJob/operator scopeAutoscaler-specificYesComparative benchmark
This studyThree synthetic enterprise domains30/60/120 sCross-domain leading signalsSLO, cost, churn, fairness450-run computational simulationSeeds, code, CSV, models, checksums

3. System Model and Problem Formulation

3.1 Domains, Events, and Workload State

Let D={o,p,s} denote the order-inventory, payment-finance, and customer-service domains. At control time t, domain d has an observed event arrival rate dt, backlog qdt, processing latency distribution Ldt, and completed throughput ydt. Events share a canonical envelope containing event identity, domain, event type, event time, ingestion time, schema version, correlation identity, partition key, run identity, and a domain-specific payload. Global ordering is neither required nor claimed; records are ordered only within the relevant partition. A correlation identifier connects events that belong to the same synthetic or authorized replay cascade.

Workload class identifies the scenario family: stable, periodic, ramp, sudden burst, cross-domain cascade, or concept drift. The system maintains causal rolling features over a five-minute history. Features include short and long event-rate windows, arrival slopes, backlog, latency percentiles, resource utilization, ready replicas, temporal encodings, capacity pressure, previous actions, and lagged cross-domain signals. Every feature has an event-time cutoff and a freshness indicator. A feature newer than the prediction origin or older than the stale threshold is rejected, preventing future leakage and unsafe use of incomplete telemetry. Section 5.4.1 details the transformation and data-quality rules.

3.2 Capacity and Decision State

For each domain, the simulated execution state includes ready replicas, fixed CPU and memory requests per replica, and an empirically calibrated safe service rate per replica. Shared CPU and memory feasibility is calculated by multiplying each candidate replica count by these fixed requests after system reservations; CPU and memory tiers cannot change during a run. A replica decision takes effect only after the modeled readiness delay. The controller maintains minimum and maximum replicas, protected domain minima, a maximum per-cycle step, minimum dwell time, asymmetric scale-out and scale-in cooldowns, and an incident or fallback flag.

The horizontal action vector is

at = {rd*(t) | d ∈ D}.

where each component is a target ready-replica count for one domain. No CPU- or memory-tier variable appears in the action space. Exactly one emulated authority controls a workload in a run: static mode has no active scaler, reactive mode uses a tuned horizontal policy emulator, and predictive mode uses the proposed horizontal controller emulator. A mode conflict invalidates the run before workload execution.

3.3 Forecast and Uncertainty

For forecast horizon h∈{30,60,120} seconds, the forecast service returns dt+h, a predictive spread dt+h, a normalized uncertainty score, the model and feature versions, inference latency, and an expiry time. Intuitively, the uncertainty buffer adds demand headroom in proportion to forecast spread: less certain forecasts produce a larger allowance, while confident forecasts remain close to the point estimate. The multiplier controls how conservative that allowance is. Formally:

defft+h=dt+h+zdt+h

The first feasible replica estimate is

rdrawt+h=ceildefft+htargetd

The raw estimate is then bounded by the protected minimum and maximum:

rd*t+h=minrdmax,maxrdmin,rdrawt+h

where target=0.65 is the provisional target occupancy. The pilot may calibrate service rates and final thresholds, but the main experiment cannot be used for tuning.

3.4 Objective and Constraints

Candidate actions are evaluated using

Ja=wslaRslaa+wcostCresourcea+wchurnPscalea+wfairIdomaina

The first term estimates service-level risk from forecast demand, backlog, and capacity. The resource term represents requested replica capacity: replica active time is translated into vCPU- and memory-hours using fixed per-replica requests. It is an accounting term, not a vertical-scaling decision variable. The remaining terms penalize repeated or opposite-direction actions and imbalance in normalized domain service satisfaction. Hard constraints take precedence over the objective: cluster budget, replica limits, protected minima, maximum step, cooldown, dwell, no scale-in during an incident, forecast validity, telemetry freshness, and single-authority enforcement. If no candidate is feasible, the guard retains the last safe allocation and records the violated constraint.

Table 2. Core notation and reference service objectives

Symbol or termMeaningUnit or reference
dtOffered event rate for domain devents/s
qdtQueue or consumer backlogevents
rdtReady replicascount
cdt,mdtCPU and memory requests per workloadvCPU, GiB
dSafe service rate per replica at target occupancyevents/s
hForecast horizon30, 60, or 120 s
dDecision-to-readiness delayseconds
Order-inventory SLOp95 latency; completion; error≤750 ms; ≥98%; ≤1%
Payment-finance SLOp95 latency; completion; error≤500 ms; ≥98%; ≤1%
Customer-service SLOp95 latency; completion; error≤1000 ms; ≥98%; ≤1%

4. Unified Orchestration Architecture

4.1 Architecture Overview

The target implementation architecture separates five responsibility layers while linking them through versioned contracts. Enterprise sources generate domain events; adapters validate the canonical envelope and schema; Kafka provides durable transport; Apicurio Registry governs Avro compatibility; and Flink performs event-time transformations, joins, windows, and domain sinks. OpenTelemetry and Prometheus would expose traces, structured context, application metrics, queue state, platform state, and controller behavior, while TimescaleDB would store aligned features, forecasts, decisions, actions, and experiment metadata. The predictive control plane converts this evidence into a guarded horizontal replica target, and a future Kubernetes execution interface would apply the authorized scale action and report readiness.

Figure 1. Logical architecture of the unified data plane and predictive control plane.

The separation between data flow and control flow provides two target safety properties. First, a forecast or controller failure should not remove the durable ingestion path. Second, control overhead should be measurable separately from event-processing latency. The predictive controller is designed as a model-agnostic layer over Kafka, Flink, and Kubernetes. In the present evaluation, these physical components are represented by deterministic workload, service-capacity, queue, readiness-delay, and controller modules; Figure 1 therefore describes the target implementation architecture rather than a deployed testbed.

Table 3. Principal architecture components

ComponentResponsibilityMain inputMain outputFailure behavior
Event adaptersNormalize and validate eventsDomain source recordsValid event or dead letterReject invalid schema; preserve evidence
Kafka and schema registryDurable transport and contract governanceVersioned eventsOrdered partition recordsRetain/replay; expose broker health
Apache FlinkEvent-time processing and stateKafka topicsDomain results and telemetryCheckpoint and recover; expose backpressure
Feature builderAlign rolling and cross-domain featuresPrometheus and audit historyFeatureWindowMark missingness; reject stale windows
Forecast servicePredict demand and uncertaintyValid FeatureWindowForecastResponseSeasonal-naive once, then fallback
Allocation optimizerScore feasible plansForecast, capacity, SLOAllocationPlanReturn infeasibility with objective terms
Policy guardEnforce safety and stabilityCandidate plan and cluster stateGuardedDecisionHold last safe plan
Predictive controllerApply and verify horizontal replica actionGuardedDecisionEmulated ActuationResultIdempotent retry, alert, and hold
Experiment orchestratorIsolate strategy and scenarioFrozen manifestRun identity and bundleAbort on authority or hash conflict

4.2 Event, Telemetry, and Evidence Contracts

The event envelope carries event_id, event_type, domain, occurred_at, ingested_at, schema_version, correlation_id, partition_key, run_id, seed, and payload. Minor schema versions must be backward compatible. Breaking changes require a new major version and an explicit migration scenario. Flink initially uses a 10-second watermark and 30-second allowed lateness. Ingestion is at least once, while duplicate domain effects are prevented using event identity. Exactly-once behavior is claimed only for paths that are actually covered by checkpointed Flink semantics.

Telemetry is sampled at a five-second cadence for application, Kafka, Flink, Kubernetes, and cluster metrics. The control interval is 15 seconds. Event-level latency uses source event time and terminal completion time; aggregated windows are aligned to UTC. Prometheus labels remain bounded: run identity is allowed only on experiment-scoped metrics, while event identity, correlation identity, user-like identifiers, payload values, and raw URLs are excluded. High-cardinality evidence is stored in bounded Parquet or trace files rather than metric labels.

Two identities create the provenance chain. A run_id binds scenario, seed, strategy, software images, schemas, SLOs, policies, and infrastructure snapshot. A decision_id binds the observation window, feature version, forecast, candidate replica targets, constraints, chosen action, actuation response, readiness delay, and observed outcome. In this study the actuation response is generated by the deterministic simulator; a future implementation would record the Kubernetes response under the same contract. The same valid inputs and frozen configuration must produce the same AllocationPlan, and a decision identity cannot be applied twice.

4.3 Predictive Decision Sequence

Each horizontal-control cycle moves through OBSERVE, VALIDATE, PREDICT, PLAN, GUARD, APPLY, VERIFY, and COOLDOWN states. OBSERVE reads workload and capacity signals. VALIDATE checks completeness, timestamps, schema identity, and scaling authority. PREDICT requests horizon estimates before a deadline. PLAN enumerates feasible replica targets. GUARD applies capacity, fairness, hysteresis, cooldown, dwell, and fallback rules. In the computational evaluation, APPLY and VERIFY are deterministic state transitions that impose the bounded readiness delay and record the modeled capacity response. In a future Kubernetes implementation, those states would use optimistic concurrency against the scale subresource and distinguish API acknowledgment from actual pod readiness. COOLDOWN prevents an unstable reversal and writes the audit record.

Figure 2. Predictive allocation control loop and fallback boundary.

The forecast response expires at valid_until. A technically successful response received too late is not valid for the current decision. Uncertainty above the frozen threshold expands the buffer, blocks scale-in, or activates fallback according to policy. This design prevents low-confidence forecasts from being treated as precise operational commands.

4.4 Horizontal Replica Control and Scope Boundary

The only active scaling loop changes horizontal replicas every 15 seconds and uses the 60-second horizon as the provisional primary prediction. Thirty- and 120-second horizons support calibration and sensitivity analysis. The provisional replica range is 2-12 per domain, maximum change is +4 or -2 replicas per cycle, scale-out cooldown is 60 seconds, scale-in cooldown is 180 seconds, and minimum dwell is 120 seconds. These values reflect asymmetric risk: late scale-out can damage an SLO quickly, while premature scale-in can recreate the same pressure and cause flapping.

CPU and memory request tiers remain fixed throughout every simulated run, so the controller has no vertical action and no slow loop. All reported comparative effects therefore arise solely from horizontal replica decisions. A separate slow right-sizing loop is reserved for future work because credible evaluation would need to model pod replacement or restart, state restoration, rollout duration, and the resulting latency penalty rather than treating a tier change as instantaneous.

Table 4. Frozen computational controller parameters

ParameterFrozen valueSimulation rationale
Fast control interval15 sMatches the controller decision cadence
Feature history5 minSupplies lags, slopes, rolling mean and SD
Forecast horizons30/60/120 sBrackets the emulated readiness delay
Primary horizon60 sFrozen controller input
Target occupancy0.65Maintains modeled headroom
Replica range2-12/domainProtected minimum and bounded action space
Maximum replica step+4/-2Faster expansion than contraction
Out/in cooldown45/240 sPilot-calibrated stability control
Telemetry stale threshold30 sTwo control cycles
Readiness delayN(52, 9^2) s; 30-80Emulates decision-to-ready latency
Vertical CPU/memory scalingExcluded; requests fixedFuture work; outside current scope

4.5 Safety, Fallback, and Governance

Safety is part of the allocation policy rather than a post-processing check. The controller uses least-privilege access: it can read approved workload state and patch only approved scale targets. Workload pods have no scaling permission. Data-plane, control-plane, observability, and workload namespaces have separate quotas, network policies, and service identities. Images are pinned by digest, containers run without unnecessary privilege, and credentials are excluded from images, manifests, model artifacts, metrics labels, and archived logs.

If telemetry is stale, the controller rejects the feature window and holds the last safe allocation. After two invalid cycles it switches to the tuned reactive fallback. If the forecast service times out, a seasonal-naive forecast may be used for one cycle before fallback. High uncertainty increases the safety buffer and prohibits scale-in. The failure campaign represents a rejected scale action as an emulated actuation error, retries it once under the same decision identity, and blocks further contraction after a readiness timeout. Capacity exhaustion activates protected domain minima, and a simulated authority conflict invalidates the run. These are policy-level results; they are not observations from a Kubernetes API.

The AI boundary is narrow. The model predicts aggregate compute demand and does not decide about individuals, payments, credit, employment, pricing, or service eligibility. Synthetic events contain no personal identifiers. Any future use of production traces requires authorization, minimization, irreversible de-identification or aggregation, controlled access, and a documented retention policy.

5. Computational Simulation Method

5.1 Research Design and Evidence Claims

The evaluation combines DSR with a deterministic, discrete-time, repeated-run computational simulation (Law, 2015). One experimental unit is a 45-minute simulated run defined by workload, strategy, and seed: 10 minutes of warm-up, 30 minutes of measurement, and five minutes of cooldown. The one-second state transition updates offered demand, ready replicas, service capacity, queue, completions, and latency samples. Control decisions occur every 15 seconds. Workload and service-noise trajectories are identical across strategies within each workload-seed block; the seed, rather than second-level observations, is the replication unit.

Three hypotheses are confirmatory within the simulation boundary. H1 tests whether predictive allocation reduces p95 and p99 latency relative to tuned reactive control during ramp, burst, and cascade workloads. H2 tests whether it reduces service-level violation rate under capacity pressure. H3 tests whether it reduces normalized resource units relative to conservative static provisioning while maintaining throughput. H4 and H5 are supportive mechanism hypotheses concerning cross-domain features, uncertainty buffering, and constrained allocation. No hypothesis is interpreted as evidence of physical Kubernetes effectiveness.

Table 5. Research question, contrast, and required evidence

RQMain contrast or testRequired evidence
RQ1Logical and computational traceabilityConfig-to-run identity and deterministic regeneration; physical actuation excluded
RQ2F0-F2 by horizon; F3 complexity-excludednRMSE, wMAPE, coverage, width, inference timing, model size
RQ3S2 vs. S1 and S0p95/p99, violations, throughput non-inferiority
RQ4Strategy x workloadRequested NCU-hours, actions, oscillation, fairness
RQ5A1-A4 and F1-F4paired mechanism deltas, oracle gap, fallback, recovery, safety

5.2 Simulation Environment and Model Identity

The campaign executed on Linux 6.18.35 (x86_64) with nine allocated AMD EPYC 9V74 virtual CPU cores and 15 GiB RAM. The analysis environment used Python 3.12.13, NumPy 2.3.5, pandas 2.2.3, SciPy 1.17.0, and scikit-learn 1.8.0. The simulator is a fluid queue and control model, not a Kubernetes emulator: it represents replica capacity, stochastic service variation, decision cadence, bounded readiness delay, queue accumulation, and constrained target application without starting containers or invoking a Kubernetes API.

The frozen configuration SHA-256 is d83c43cd9e74a16ad060471b7230a940305804ad1c133f9105f5678506fd5bb7. Each run is reconstructable from the configuration, workload identifier, strategy or variant, and numeric seed. The primary 60-second F2 model bundle occupies 3,106,313 bytes. Wall-clock inference timing is host-specific and is reported only as implementation context; it is not included in simulated resource consumption. Simulation verification follows trace checks, invariant checks, independent rerun, and output checksum comparison (Sargent, 2013).

5.3 Workload Generator and Domains

The generator produces deterministic enterprise-like rate and service trajectories from the frozen configuration and a seed. Reference base rates are 600 events/s for order-inventory, 420 events/s for payment-finance, and 180 events/s for customer service. Per-replica safe service rates are 400, 300, and 140 events/s, respectively. Shared and domain-specific autoregressive noise perturb arrival and service rates while preserving identical trajectories across paired strategy runs.

The workload families are:

  • W0 Stable: 1.0× base load for the 30-minute measurement phase.
  • W1 Periodic: a sinusoidal pattern from 0.7× to 1.5× base with a five-minute period.
  • W2 Ramp: 0.6× to 2.0× over 10 minutes, followed by a 10-minute plateau and a controlled descent.
  • W3 Sudden burst: 1.0× to 3.0× within 15 seconds, held for five minutes and repeated twice.
  • W4 Cross-domain cascade: an order burst followed by payment at +30 seconds, inventory at +45 seconds, and customer service at +120 seconds.
  • W5 Concept drift: a 50% magnitude increase and altered cascade lags after minute 15.

The simulator uses a one-second fluid update rather than materializing every event. Completed volume equals the minimum of queued-plus-arriving work and available replica capacity. Queue delay is combined with a utilization-pressure term and eight deterministic log-normal latency samples per domain-second; completion volume supplies the sample weights. This abstraction preserves overload, backlog, service-level violations, and recovery dynamics, but it does not model event schemas, payload serialization, partitions, checkpoints, or network packets.

Table 6. Frozen workload profiles

IDShapePrincipal stressPrimary purpose
W0Constant 1.0×Steady overheadEfficiency and controller overhead
W10.7×–1.5× sinusoidRepeated peaksSeasonality and stable scale cycles
W20.6×–2.0× rampGradual pressureLead time and conservative scale-in
W33.0× burst in ≤15 sAbrupt overloadTail latency, backlog, uncertainty
W4Staggered domain cascadeCorrelated propagationValue of cross-domain leading signals
W5Magnitude and lag shiftDistribution changeRobustness, uncertainty, fallback

5.4 Forecasting Pipeline, Model Selection, and Drift Control

5.4.1 Feature Engineering and Data-Quality Handling

At each forecast origin t, the feature builder follows a causal six-stage sequence: raw one-second domain and service-state trajectories -> event-time alignment -> data-quality screening -> lag, rolling, and slope calculation -> cross-domain joining -> the F2 input matrix. All source values are truncated at t before transformation. For each domain, the builder selects rate lags at 0, 15, 30, 60, 120, and 300 seconds for all three domains; computes 60-second rolling means and standard deviations; derives 30- and 60-second arrival-rate slopes; and appends current backlog, ready replicas, capacity pressure, temporal encodings, and freshness metadata. Local and leading cross-domain features are joined only on the common event-time cutoff, so no value after the prediction origin can enter a training or inference row.

The executed synthetic traces are generated on a complete one-second grid, so the reported training, validation, and locked-test matrices require no statistical imputation. Rows lacking the full 300-second causal history are excluded until warm-up is complete. A non-finite, negative, stale, or otherwise incomplete required value invalidates the row; the pipeline does not use future-looking interpolation, bidirectional filling, or global mean/median replacement. Valid extreme values created by ramp, burst, cascade, and drift scenarios are retained rather than winsorized because they are experimental signals, not nuisance outliers. Consequently, data-quality handling cannot smooth away the workload stresses that the forecast and controller are intended to face.

5.4.2 Forecast Models, Selection, and Drift Control

Forecast comparison includes a current-value seasonal-naive proxy (F0), a 60-second trailing mean (F1), and histogram-based gradient-boosted trees (F2). F2 uses rate lags of 0, 15, 30, 60, 120, and 300 seconds for all domains, 30- and 60-second slopes, and 60-second rolling means and standard deviations. Absolute run time and future outcomes are excluded to prevent scheduled-burst leakage. The prespecified complexity gate required at least 100 independently seeded training traces before a neural sequence candidate could be evaluated. The W0-W4 training split contained 60 traces (12 seeds x five workload profiles), so F3 was excluded.

Independent synthetic traces for W0-W4 are split by seed: 4101-4112 for training, 4113-4115 for validation, and 4116-4118 for the locked forecast test. Separate mean, 0.10-quantile, and 0.90-quantile F2 models are fit for each domain and for 30-, 60-, and 120-second horizons. The 60-second horizon is the primary controller input. Local-only F2 models are trained separately for the A1 ablation. Selection uses test-independent validation comparisons of nRMSE, wMAPE, empirical interval coverage, and interval width.

Forecast accuracy is evaluated with nRMSE and wMAPE because unweighted percentage error is unstable near zero (Hyndman & Koehler, 2006). The 0.10-0.90 interval is evaluated using empirical coverage and mean width. F2 is frozen before the 450-run campaign because it yields the lowest validation errors across the three horizons. Models are not retrained during a run or after main outcomes are observed.

Latency, service-level outcomes, allocation actions, future demand, and future backlog are unavailable as features. W5 is excluded from model training and changes both cascade magnitude and inter-domain lags after minute 15. A relative forecast-error detector may activate the reactive fallback after two cycles above the frozen 0.40 threshold; failure to activate is retained as a robustness outcome rather than corrected after inspection.

5.5 Treatments and Fair Comparison

S0 is conservative static allocation, frozen at 7, 6, and 5 replicas for order-inventory, payment-finance, and customer service from the pilot 95th-percentile rates at 0.70 occupancy. S1 emulates tuned reactive horizontal scaling from current 30-second mean rate, occupancy, and a 30-second backlog-drain target. S2 emulates predictive horizontal scaling using the F2 60-second forecast, half of the 0.90-quantile uncertainty gap, queue pressure, a constrained allocator, protected minima, and fallback. S1 and S2 share replica bounds of 2-12 per domain, a 24-replica total budget, 0.65 target occupancy, maximum steps of +4/-2, the same bounded readiness-delay function, and fixed CPU and memory requests per replica.

Only P01-P05 pilot seeds across W0, W2, W3, and W4 are used to calibrate the static plan, reactive thresholds, backlog drain, uncertainty weight, forecast smoothing, and cooldowns. The final scale-out and scale-in cooldowns are 45 and 240 seconds; readiness delay is normally centered at 52 seconds with 9-second standard deviation and clipped to 30-80 seconds. A predictive scale-in deadband of two replicas limits churn. These values and the main seed register are frozen before confirmatory outputs are opened.

Table 7. Treatment fairness controls

DimensionS0 StaticS1 ReactiveS2 Predictive
Scaling authorityNoneReactive policy emulatorPredictive policy emulator
Workload and seedIdenticalIdenticalIdentical
Replica and total limitsSame boundsSame boundsSame bounds
Service modelSame rates/noiseSame rates/noiseSame rates/noise
SLO/readiness functionSameSameSame
Decision inputFrozen planCurrent rate/backlogForecast, uncertainty, queue
SafetyStatic boundsBounds/cooldownGuard, minima, fallback

5.6 Pilot Execution and Parameter Freeze

The pilot is a calibration and model-verification stage, not a smaller confirmatory experiment. Its 60 runs verify queue and capacity invariants, establish the static plan, tune S1 using the same evidence budget as S2, freeze the selected forecast and controller parameters, and confirm that all run identities and outputs can be regenerated. Pilot outcomes are excluded from confirmatory effect estimates.

The freeze register binds the configuration hash, scenario definitions, S0 plan, S1 thresholds, F2 model contract, S2 uncertainty and stability parameters, resource-index formula, seed lists, primary contrasts, and claim gates. Main-run results are never used to retune a treatment. Two later changes were permitted as design corrections before the final campaign: preventing simultaneous pending actions for one domain and pairing ablation seeds with their full-S2 comparator; the complete final campaign was then rerun from the frozen configuration.

Pilot data support calibration only. Confirmatory interpretation uses 15 seeds per W0-W4 strategy cell, with W5 retained as a separate drift evaluation. All simulation-level claims are conditioned on the frozen mathematical model and are explicitly separated from claims that would require a physical stream-processing cluster.

5.7 Run Lifecycle, Pairing, and Campaign

Each run begins from a fresh in-memory state using a frozen workload-strategy-seed identity. A 10-minute warm-up establishes base rate and initial replicas; the next 30 minutes are measured, followed by five minutes of reduced offered load and queue drain. The simulator records run-level latency samples, violations, completed and offered volume, requested compute units, replica trajectories, actions, oscillations, fairness, fallback behavior, and safety invariants.

Figure 3. Controlled computational-simulation run lifecycle.

A block is defined by workload and seed. All strategies within a block receive byte-equivalent offered-load, service-factor, and latency-randomness arrays. Because every run is a deterministic isolated function call with no shared mutable state, execution order cannot introduce day or carryover effects and Latin-square ordering is unnecessary. Ablation uses confirmatory seeds 1001-1010 so every variant can be paired with the full S2 result.

The executed campaign contains 60 pilot runs, 225 confirmatory runs, 45 drift runs, 80 ablation runs, and 40 failure-injection runs, totaling 450 attempts. Fifteen independent seeds are used per confirmatory cell and ten paired seeds per ablation or failure condition. This fixed count was selected before final outcomes were inspected and exceeds the prespecified minimum of 15 seeds per confirmatory cell (Lakens, 2022).

Table 8. Executed computational simulation campaign

StageWorkloadsConditionsReplicationPlanned runs
P0 PilotW0, W2, W3, W4S0, S1, S25 seeds60
E1 ConfirmatoryW0–W4S0, S1, S215 seeds225
E2 DriftW5S0, S1, S215 seeds45
E3 AblationW3, W4A1–A410 seeds80
E4 FailureW4 plus faultS2/fallback, four faults10 seeds40

5.8 Measures, Data Quality, and Statistical Analysis

Primary outcomes are weighted run-level p95 latency, service-level violation rate, and normalized requested horizontal replica capacity per million completed events. One normalized compute unit (NCU) is defined as one vCPU-hour plus one quarter of a memory-GB-hour, calculated from replica active time and fixed per-replica requests. Throughput completion ratio is a non-inferiority guardrail with a -5% margin. Secondary outcomes are p99 latency, requested vCPU- and memory-GB-hours, action count, oscillation, Jain fairness, fallback time, recovery time, and protected-minimum compliance.

The simulation reports requested capacity only; CPU usage, memory working set, network, disk, and energy are not modeled. Oscillation is an opposite-direction action within two cooldown windows. Fairness uses Jain’s index over domain SLO satisfaction (Jain, 1991). Percentages retain their denominators, and every primary contrast names the comparator, workload scope, paired effect, 95% bootstrap confidence interval, and multiplicity status.

A run is invalid if its identity is duplicated, any state becomes non-finite, completed work exceeds offered-plus-queued work, a replica leaves the frozen domain bounds, the total applied plan exceeds the active budget, or output columns are incomplete. No final-campaign run triggered these conditions. Treatment-induced queues, violations, and fault responses remain valid outcomes. Pilot rows are retained but excluded from confirmatory effects.

Primary contrasts use paired percentage changes within workload-seed blocks. The reported effect is the median paired change with a 5,000-resample percentile bootstrap confidence interval. Two-sided Wilcoxon signed-rank tests provide p-values for the latency, violation, and resource contrasts, and Holm correction controls that three-test confirmatory family (Holm, 1979). Throughput is assessed as a paired difference in completion ratio; its lower 95% bound must remain above -0.05.

Workload-specific paired summaries expose heterogeneity that an aggregate median can hide. The interpretation emphasizes practical thresholds in addition to adjusted p-values. A complete rerun from the same configuration and seeds must reproduce the primary run and contrast CSV files byte for byte. This procedure treats seeds as independent realizations within the model while recognizing that repeated simulation does not establish external validity (Jain, 1991; Kalibera & Jones, 2013; Law, 2015).

Table 9. Operational outcome definitions

OutcomeSimulation definitionReporting
End-to-end latencyQueue delay + utilization pressure + weighted log-normal samplesp95/p99 ms by domain and run
SLA violationWeighted latency samples above domain thresholdRate per run
ThroughputCompleted fluid volume / offered volumeCompletion ratio
BacklogUnfinished volume after one-second service updateEvents
Requested resourceReplica request x active timevCPU-h, GB-h, NCU-h/million
Ready-capacity responseApplied target time minus detected pressure onsetSeconds; diagnostic only
OscillationOpposite-direction actions within two cooldown windowsCount/run
FairnessJain index over domain SLO satisfaction0-1

5.9 Ablation, Failure Injection, and Reproducibility

A1 replaces the full forecast with separately trained local-domain models; A2 removes the uncertainty buffer; A3 removes constrained stability behavior and uses a greedy forecast target; and A4 uses demand 60 seconds in the future as an oracle upper bound. A1-A4 are evaluated with ten paired W3 and W4 seeds. The oracle is diagnostic and is never treated as an operational competitor.

Four control-plane failures are emulated at a seeded point between minutes 10 and 20 of measurement. F1 withholds forecasts for 120 seconds and activates reactive fallback. F2 withholds fresh telemetry and requires a hold decision. F3 rejects the first eligible scale patch. F4 reduces the total replica budget from 24 to 16 for 180 seconds. Acceptance requires F1 fallback within 30 seconds, no decision from stale telemetry, no duplicate plan after rejection, and no domain below two replicas during pressure.

The reproducibility bundle contains the simulation source, frozen JSON configuration, seed register, primary-horizon model artifact, 450-row run dataset, forecast evaluation, paired contrasts, joint-gate record, figures, and SHA-256 manifest. R1 regenerates all outputs from code and configuration. Physical workload replay (R2) and full Kafka/Flink/Kubernetes replication (R3) are not part of the present evidence and remain future validation levels.

6. Results

6.1 Run Accounting and Data Quality

All 450 scheduled simulation runs completed and were retained: 60 pilot, 225 confirmatory, 45 drift, 80 ablation, and 40 failure-injection runs. No run violated the numerical, identity, replica-bound, or capacity invariants, and no replacement was required. Pilot observations were used only for calibration. The confirmatory dataset contained 75 complete workload-seed blocks (five workloads x 15 seeds) and 2,700,000 weighted domain-second latency samples from the 30-minute measurement phases.

Table 10. Simulation run accounting

StageScheduledAttemptedValidInvalidSafety stopReplacementAnalyzed
P0 Pilot606060000Calibration only
E1 Confirmatory225225225000225
E2 Drift45454500045
E3 Ablation80808000080
E4 Failure40404000040

6.2 RQ1: Integration Feasibility

Every scheduled run was linked to a unique workload, strategy or variant, and seed, and every result row was regenerated from the same configuration hash. The second full execution produced byte-identical run_results.csv and primary_contrasts.csv files; the final run-results SHA-256 is 8ac565365f6885c2995578e3fdd3d4f8bc6946b5bc82198ce74d66237ef6cf52. This establishes computational traceability from configuration to aggregate outcome.

RQ1 is therefore supported only at the logical and computational-integration level. The simulation does not supply API acknowledgments, actual pod readiness, broker offsets, Flink checkpoints, or event-to-decision traces from deployed components. Those architecture acceptance criteria require a later R2/R3 implementation and cannot be inferred from the 450 simulated runs.

6.3 RQ2: Forecast Accuracy, Calibration, and Overhead

On the locked W0-W4 forecast test, F2 produced nRMSE values of 0.145, 0.240, and 0.330 at 30, 60, and 120 seconds; corresponding wMAPE values were 0.053, 0.096, and 0.154. At the primary 60-second horizon, F2 improved nRMSE by 28.2% relative to F0 (0.240 versus 0.335). Its 0.10-0.90 interval covered 75.2% of targets with mean width 133.0 events/s, indicating 4.8 percentage points of undercoverage relative to an 80% nominal interval. Batch inference required 0.018 ms per control row on the execution host, and the nine-model primary-horizon bundle occupied 3.0 MiB. No container-level CPU or memory overhead was measured.

Figure 4. Forecast nRMSE by candidate model and horizon on the locked synthetic test set.

Table 11. Locked forecast test results

ModelHorizonDomain/workloadnRMSEwMAPECoverageMean widthInference msMemory MiB
F0 Current-value proxy30/60/120Pooled W0-W4 test0.217/0.335/0.4950.085/0.144/0.248N/AN/A<0.001N/A
F1 Trailing mean30/60/120Pooled W0-W4 test0.310/0.403/0.5360.136/0.189/0.281N/AN/A<0.001N/A
F2 Boosted trees30/60/120Pooled W0-W4 test0.145/0.240/0.3300.053/0.096/0.1540.754/0.752/0.74677.7/133.0/286.30.040/0.018/0.0123.0 (60-s bundle)
F3 Sequence modelNot runComplexity gateN/AN/AN/AN/AN/AN/A

6.4 RQ3: End-to-End Performance

Across 45 paired W2-W4 blocks, S2 reduced p95 latency by a median 49.5% relative to S1 (95% CI [10.2%, 70.4%]; Holm-adjusted p < .001). The pooled descriptive medians were 669.8 ms for S2 and 8,932.9 ms for S1. S2 reduced the paired service-level violation rate by 46.4% (95% CI [43.7%, 49.8%]; adjusted p < .001), with descriptive medians of 5.00% and 12.97%. Effects were heterogeneous: p95 reductions were 8.4% for W2, 60.7% for W3, and 91.9% for W4, while W0 and W1 reductions were 11.8% and 33.5%.

Relative to S0 across W0-W4, S2 reduced normalized requested resources by a paired median 29.8% (95% CI [26.9%, 36.5%]; adjusted p < .001). Descriptive medians were 2.684 and 3.307 NCU-hours per million completed events. S2 used more normalized capacity than S1 in W0-W3 but 5.8% less in W4, showing that its efficiency claim is against conservative static provisioning, not against reactive control in every workload. Median completion ratios were 1.000 for all three strategies; the lower paired difference bound was 0.000, above the -0.05 non-inferiority margin.

Figure 5. Median run-level p95 latency by workload and allocation strategy.

Table 12. Primary simulation contrasts

Outcome and workloadComparatorS2 estimateComparator estimatePaired effect95% CIAdjusted pGate
p95 latency, W2-W4S1669.8 ms8,932.9 ms49.5% reduction10.2%-70.4%<.001G1 PASS
SLA violation, W2-W4S15.00%12.97%46.4% reduction43.7%-49.8%<.001G2 PASS
Normalized NCU, W0-W4S02.6843.30729.8% reduction26.9%-36.5%<.001G3 PASS
Completion ratio, W0-W4S0/S11.0001.000/1.000Difference 0.0000.000-0.000N/AG4 PASS

6.5 RQ4: Efficiency, Stability, and Fairness

Across W0-W4, S2 issued a median 16 scaling actions per run versus 22 for S1. Median opposite-direction oscillations were 7 and 9, respectively. S2 maintained a median 12.79 total replicas during W2 and 15.44 during W3, compared with 12.04 and 15.04 under S1; predictive benefit therefore arose from earlier placement and workload-specific allocation rather than uniformly higher capacity.

The median Jain fairness index was 0.999991 for S2, 0.999631 for S1, and 0.999999 for S0. No run breached the two-replica protected minimum, and the final application rule prevented total ready replicas from exceeding 24. When drift runs were included, median oscillations were nine for both S2 and S1; consequently, the stability gate passed as non-worsening rather than as a claim that predictive control eliminated churn.

Figure 6. Requested-resource and p95-latency trade-off across confirmatory simulation runs.

6.6 RQ5: Mechanisms, Drift, and Failure Recovery

Mechanism effects depended on workload. Under W4 cascades, replacing cross-domain forecasting with local-only models increased paired p95 latency by 123.7% (95% CI [12.6%, 211.6%]) and violations by 64.7% (95% CI [13.1%, 112.1%]). Under simultaneous W3 bursts, however, A1 reduced p95 by 50.8%, showing that cross-domain inputs can add noise when no leading relation exists. Removing the uncertainty buffer increased p95 by 93.5% across W3-W4, while the greedy A3 policy increased p95 by 226.0%, violations by 53.3%, and oscillations by nine despite using 15.3% fewer resources. The oracle reduced p95 by 71.7%, leaving a material forecast-limited gap.

Drift reversal (W5). S2 cut median p95 latency by 64.5% versus S1, yet its 4,638.3 ms result was 845.3% higher than S0’s 490.7 ms. Its 13.18% violation rate was 32.9% lower than S1 but 841.4% higher than S0’s 1.40%. Because the frozen 0.40 error detector did not trigger fallback, W5 demonstrates unresolved drift sensitivity, not drift robustness.

Figure 7. W5 concept-drift reversal: predictive control outperformed reactive control but remained substantially worse than static allocation.

For failure recovery, F1 activated reactive fallback after a median five seconds, and queues recovered within 20 seconds after fault clearance. F2 produced zero stale actions; F3 produced one rejected patch and zero duplicate plans; F4 preserved the two-replica domain minimum and recovered within 26 seconds.

6.7 Joint Success Decision

Table 13. Joint success gates for the computational simulation

GateThresholdObserved evidenceStatusPermitted interpretation
G1 Latency>=15% p95 improvement vs S149.5%; CI 10.2%-70.4%PASSSimulation-level latency benefit
G2 SLA>=20% violation reduction vs S146.4%; CI 43.7%-49.8%PASSSimulation-level SLO benefit
G3 Efficiency>=10% NCU reduction vs S029.8%; CI 26.9%-36.5%PASSLower requested capacity than static
G4 ThroughputLower bound above -5%Lower paired difference 0.000PASSCompletion non-inferior in model
G5 StabilityNo worse oscillation; bounds hold9 vs 9; domain >=2; total <=24PASSNon-worsening modeled stability
G6 ReproducibilityPrimary outputs regenerateTwo primary CSVs byte-identicalPASSComputational reproduction only

G1-G6 all passed within the frozen simulation: latency, violation, efficiency, throughput, stability, and computational reproducibility met their thresholds. The permitted conclusion is correspondingly bounded: S2 outperformed the tuned reactive policy on pressure-sensitive simulated outcomes and used less requested capacity than S0 without sacrificing simulated completion. The gate result does not authorize a claim of physical deployment readiness, and the adverse W5 outcome prevents a general claim of robustness to concept drift.

7. Discussion

7.1 Interpretation Framework

The evidence answers RQ1-RQ5 at two distinct levels. The architecture specifies an auditable event-to-horizontal-allocation loop, while the experiment validates only its deterministic computational representation. F2 supplies useful but undercovered forecasts; S2 improves pressure-sensitive simulated service outcomes; its requested-capacity advantage is relative to static allocation; and its mechanisms are workload-dependent. This separation prevents favorable simulated latency from being mistaken for proof of a working Kubernetes control plane.

The six simulation gates passed, but the operating region matters. S2 produced modest p95 improvement during the gradual W2 ramp, larger improvement during simultaneous W3 bursts, and the largest benefit during W4 cascades. W0 remained a case in which S1 used less capacity and already maintained a low violation rate. The strongest interpretation is therefore not universal predictive superiority; it is that lead time and guarded allocation matter most when reactive readiness delay allows pressure to accumulate.

7.2 Mechanisms and Trade-Offs

Ablations separate three mechanisms. Cross-domain features supplied substantial leading information in W4, yet harmed W3 where domains rose together rather than sequentially. The uncertainty buffer protected tail outcomes with little median resource penalty. The constrained policy guard was especially important: A3 saved requested capacity but paid for it with markedly higher tail latency, violations, and oscillation. These results show that forecast accuracy alone is insufficient; the mapping from prediction to bounded action is part of the contribution.

The W5 result exposes a mechanism that erased benefit. Although S2 remained better than S1, changed magnitude and lags produced multi-second p95 latency, and the fixed error threshold failed to invoke fallback. The oracle gap likewise shows that considerable benefit remains prediction-limited. Any future deployment design therefore needs calibrated online drift detection, conformal interval correction, or a safer uncertainty trigger; those mechanisms must be evaluated prospectively rather than tuned to the observed W5 trace.

Resource interpretation is deliberately narrow. The NCU endpoint measures requested replica capacity, not observed CPU cycles, memory working set, power, or cloud cost. S2 reduced NCU use relative to S0 in every W0-W4 workload, but it used more than S1 in W0-W3. The result supports a static-efficiency claim and a pressure-performance claim; it does not support a claim that predictive allocation is always the least-resource policy.

7.3 Operational Implications

The simulation suggests that a future deployment should target operations with reliable domain relations, measurable readiness delay, bounded schemas, synchronized timestamps, and observable backlog. Cross-domain inputs should be enabled only where an upstream-to-downstream relation has been documented and monitored. The W3/W4 contrast demonstrates why one global feature policy is unsafe: a signal that is valuable in a cascade can be distracting under simultaneous demand.

A physical evaluation should begin in observe-only mode, then proceed to shadow allocation before any controller receives patch authority. The simulation-derived parameters are starting hypotheses, not production defaults. Freshness, interval calibration, capacity, provenance, protected minima, maximum steps, cooldowns, readiness verification, and fault fallback must pass on the actual cluster before comparative execution.

The present contribution is limited to horizontal replica scaling. Future work may add a separate slow CPU and memory right-sizing loop, but its evaluation must explicitly represent pod replacement or container restart, rollout and readiness delay, state restoration, Flink state movement, checkpoint recovery, and the resulting latency penalty. These effects should first be tested in an extended simulator and then in an observe-only or shadow-mode physical study; they cannot be inferred from the current NCU accounting model.

7.4 Contribution to Software Systems Research

The architecture contributes a contract-centered view of predictive horizontal autoscaling whose unit of analysis is the path from event observation to a guarded replica target. The computational artifact instantiates FeatureWindow-like histories, forecast distributions, bounded targets, readiness delay, fallback, and outcome accounting as explicit state transitions. It therefore supplies an executable theory of the target horizontal control loop while leaving broker, stream-engine, Kubernetes, and vertical-scaling integration for subsequent engineering validation.

The evaluation contributes claim discipline. Forecast accuracy is necessary but insufficient: G1-G6 jointly require latency, service-level, resource, throughput, stability, and reproduction evidence. Paired seeds, pilot-only tuning, ablations, fault emulation, visible adverse drift, and byte-identical reruns reduce the opportunity to select only favorable outputs. The same discipline also blocks overreach by separating simulation reproducibility from physical-system reproducibility.

7.5 Boundary Conditions

The conclusions apply only to the frozen discrete-time model, its three synthetic domains, service-rate assumptions, queue equation, log-normal latency sampling, readiness-delay distribution, controller policies, SLOs, and short horizons. They do not establish performance for any Kubernetes cluster, Kafka/Flink topology, stateful operator, storage system, network, cloud, or production trace. The deterministic model improves reproducibility but cannot represent every business dependency, payload, partition, checkpoint, or failure interaction.

The architecture is model-agnostic only at the forecast contract. The observed outcome depends on the synthetic training distribution, feature availability, F2 implementation, controller translation, and capacity model. F3 was not evaluated, and F2 interval coverage remained below nominal. A later study should compare alternative forecast families and calibrated uncertainty on real or authorized replay traces without changing the primary decision gates after outcomes are visible.

8. Validity, Governance, and Reproducibility

8.1 Threats to Validity

Construct validity is limited by the fluid queue abstraction and synthetic latency distribution. The model represents demand, capacity, backlog, readiness delay, and service-level outcomes, but not event-time watermarks, partition skew, serialization, checkpoint recovery, packet delay, garbage collection, or scheduler contention. Multiple workload shapes and domains improve coverage, yet they do not make the model equivalent to an enterprise deployment (Sargent, 2013).

Internal validity is strengthened by paired workload, service, and latency trajectories; fresh state per run; pilot-only calibration; frozen main seeds; and deterministic capacity constraints. Execution order, noisy neighbors, and day effects are absent by construction rather than controlled experimentally. This increases causal isolation inside the model but may overstate the regularity obtainable on physical infrastructure.

Conclusion validity rests on 15 independent model seeds per confirmatory cell, paired effects, 5,000-resample confidence intervals, Wilcoxon tests, Holm correction, practical thresholds, and workload-specific sensitivity. The resulting intervals quantify Monte Carlo variability under the frozen parameterization; they do not include structural uncertainty about whether the model is correct.

External validity is the principal limitation because no container, broker, stream processor, autoscaler, or Kubernetes API was executed. Implementation validity is likewise provisional: S1 and S2 are policy emulations sharing the same service and readiness equations, not independently engineered controllers. Physical R2/R3 replication must test scheduling, cold starts, state restore, telemetry loss, API conflicts, and real resource usage before operational recommendations are justified.

Measurement validity is exact with respect to the implemented equations but dependent on their assumptions. Weighted latency samples approximate an event distribution; NCU measures requested capacity only; completion ratio contains no packet or event loss unless backlog remains; and batch inference timing is host-specific. Source code, definitions, and checksum-bound outputs make these choices auditable rather than eliminating their limitations.

Table 14. Simulation validity threat register

TypePrincipal threatMitigation / residual boundary
ConstructFluid queue and sampled latency simplify stream processingMultiple shapes/domains; equations disclosed; physical semantics remain untested
InternalParameter or implementation biasPaired seeds, pilot-only freeze, reset state, shared service/readiness functions
ConclusionMonte Carlo uncertainty and multiple testing15 seeds/cell, paired bootstrap, Wilcoxon, Holm, practical gates
ExternalNo physical cluster or production traceClaims restricted to simulation; R2/R3 replication required
ImplementationS1/S2 are policy emulationsCommon limits and code path; no deployment claim
MeasurementNCU and latency depend on formulasDefinitions, code, model artifact, CSV, and checksums released

8.2 Security, Privacy, and Ethical Boundary

The executed campaign uses synthetic numeric trajectories only. It contains no production logs, credentials, personal identifiers, payments, customer records, or external URLs. The model predicts aggregate workload rates and never makes a decision about an individual. The main ethical obligation is transparent labeling: simulated outcomes must not be represented as observations from a real enterprise system.

The F2 models predict aggregate domain demand. Environmental impact and financial cost were not measured, and model inference timing is disclosed only for the execution host. Any generative-AI assistance used in manuscript preparation must follow the journal’s current disclosure policy; AI cannot be an author. The simulation evidence is generated by the disclosed deterministic code and seeds, not by invented narrative values.

8.3 Reproducibility and Artifact Availability

The reproducibility package contains the Python simulator, frozen JSON configuration, seed lists, F2 model artifact, forecast evaluation CSV, 450-run result CSV, primary contrast CSV, gate and freeze records, three result figures with TIF exports, and SHA-256 manifest. The configuration hash binds the workload, controller, campaign, and claim thresholds to the reported outputs.

R1 was executed twice from the frozen configuration. The two run-result and primary-contrast CSV files were byte-identical, demonstrating deterministic computational reproduction. The complete checksum-bound package accompanies the manuscript as supplementary material. R2 replay against implemented services and R3 full cluster replication are explicitly future work.

9. Conclusion

This study specifies a unified real-time orchestration architecture that links cross-domain workload telemetry, short-horizon forecasting, constrained horizontal replica allocation, fallback, and decision provenance. The evidence comes exclusively from a 450-run deterministic computational simulation comparing conservative static, tuned reactive, and predictive strategies across stable, periodic, ramp, burst, cascade, and drift conditions with paired seeds and frozen claim gates.

Within the simulation, S2 reduced paired p95 latency by 49.5% and SLA violations by 46.4% relative to S1 across W2-W4, while reducing normalized requested replica capacity by 29.8% relative to S0 and preserving simulated completion throughput. Cross-domain features were valuable for cascades but counterproductive for simultaneous bursts; the uncertainty buffer and constrained allocator protected simulated service outcomes; and concept drift remained an unresolved weakness because the frozen detector did not invoke fallback. All six simulation gates passed and primary files reproduced byte for byte. These findings support the architecture as a testable horizontal-scaling design and justify later physical validation; they do not establish production readiness, vertical-scaling effectiveness, or Kubernetes deployment performance.

References

Ahmad, H., Treude, C., Wagner, M., & Szabo, C. (2024). Smart HPA: A resource-efficient horizontal pod auto-scaler for microservice architectures. In Proceedings of the IEEE 21st International Conference on Software Architecture (pp. 46–57). IEEE. https://doi.org/10.1109/ICSA59870.2024.00013

Akidau, T., Bradshaw, R., Chambers, C., Chernyak, S., Fernández-Moctezuma, R. J., Lax, R., McVeety, S., Mills, D., Perry, F., Schmidt, E., & Whittle, S. (2015). The Dataflow model: A practical approach to balancing correctness, latency, and cost in massive-scale, unbounded, out-of-order data processing. Proceedings of the VLDB Endowment, 8(12), 1792–1803. https://doi.org/10.14778/2824032.2824076

Al-Dhuraibi, Y., Paraiso, F., Djarallah, N., & Merle, P. (2018). Elasticity in cloud computing: State of the art and research challenges. IEEE Transactions on Services Computing, 11(2), 430–447. https://doi.org/10.1109/TSC.2017.2711009

Burns, B., Grant, B., Oppenheimer, D., Brewer, E., & Wilkes, J. (2016). Borg, Omega, and Kubernetes. ACM Queue, 14(1), 70–93. https://doi.org/10.1145/2898442.2898444

Carbone, P., Katsifodimos, A., Ewen, S., Markl, V., Haridi, S., & Tzoumas, K. (2015). Apache Flink: Stream and batch processing in a single engine. IEEE Data Engineering Bulletin, 38(4), 28–38.

Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). ACM. https://doi.org/10.1145/2939672.2939785

Cortez, E., Bonde, A., Muzio, A., Russinovich, M., Fontoura, M., & Bianchini, R. (2017). Resource Central: Understanding and predicting workloads for improved resource management in large cloud platforms. In Proceedings of the 26th Symposium on Operating Systems Principles (pp. 153–167). ACM. https://doi.org/10.1145/3132747.3132772

Floratou, A., Agrawal, A., Graham, B., Rao, S., & Ramasamy, K. (2017). Dhalion: Self-regulating stream processing in Heron. Proceedings of the VLDB Endowment, 10(12), 1825–1836. https://doi.org/10.14778/3137765.3137786

Gedik, B., Schneider, S., Hirzel, M., & Wu, K.-L. (2014). Elastic scaling for data stream processing. IEEE Transactions on Parallel and Distributed Systems, 25(6), 1447–1463. https://doi.org/10.1109/TPDS.2013.295

Hevner, A. R., March, S. T., Park, J., & Ram, S. (2004). Design science in information systems research. MIS Quarterly, 28(1), 75–105. https://doi.org/10.2307/25148625

Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735–1780. https://doi.org/10.1162/neco.1997.9.8.1735

Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2), 65–70.

Hossen, M. R., Islam, M. A., & Ahmed, K. (2022). Practical efficient microservice autoscaling with QoS assurance. In Proceedings of the 31st International Symposium on High-Performance Parallel and Distributed Computing (pp. 83–94). ACM. https://doi.org/10.1145/3502181.3531460

Hyndman, R. J., & Koehler, A. B. (2006). Another look at measures of forecast accuracy. International Journal of Forecasting, 22(4), 679–688. https://doi.org/10.1016/j.ijforecast.2006.03.001

Jain, R. (1991). The art of computer systems performance analysis: Techniques for experimental design, measurement, simulation, and modeling. Wiley.

Kalibera, T., & Jones, R. (2013). Rigorous benchmarking in reasonable time. In Proceedings of the 2013 International Symposium on Memory Management (pp. 63–74). ACM. https://doi.org/10.1145/2464157.2464160

Kreps, J., Narkhede, N., & Rao, J. (2011). Kafka: A distributed messaging system for log processing. In Proceedings of the NetDB Workshop (pp. 1–7).

Lakens, D. (2022). Sample size justification. Collabra: Psychology, 8(1), Article 33267. https://doi.org/10.1525/collabra.33267

Law, A. M. (2015). Simulation modeling and analysis (5th ed.). McGraw-Hill Education.

Liu, X., & Buyya, R. (2020). Resource management and scheduling in distributed stream processing systems: A taxonomy, review, and future directions. ACM Computing Surveys, 53(3), Article 50, 1–41. https://doi.org/10.1145/3355399

Lohrmann, B., Janacik, P., & Kao, O. (2015). Elastic stream processing with latency guarantees. In 2015 IEEE 35th International Conference on Distributed Computing Systems (pp. 399–410). IEEE. https://doi.org/10.1109/ICDCS.2015.49 

Lorido-Botran, T., Miguel-Alonso, J., & Lozano, J. A. (2014). A review of auto-scaling techniques for elastic applications in cloud environments. Journal of Grid Computing, 12, 559–592. https://doi.org/10.1007/s10723-014-9314-7

Kusumba, S. (2023). A Unified Data Strategy and Architecture for Financial Mastery: AI, Cloud, and Business Intelligence in Healthcare. International Journal of Computer Technology and Electronics Communication, 6(3), 6974-6981.

Gangina, P. (2024). Generative AI integration patterns in enterprise microservices ecosystems. International Journal of Science, Research and Technology, 7(6), 13153-13165.

Luo, S., Xu, H., Ye, K., Xu, G., Zhang, L., Yang, G., & Xu, C. (2022). The power of prediction: Microservice auto scaling via workload learning. In Proceedings of the 13th Symposium on Cloud Computing (pp. 355–369). ACM. https://doi.org/10.1145/3542929.3563477

Mao, H., Alizadeh, M., Menache, I., & Kandula, S. (2016). Resource management with deep reinforcement learning. In Proceedings of the 15th ACM Workshop on Hot Topics in Networks (pp. 50–56). ACM. https://doi.org/10.1145/3005745.3005750

Montgomery, D. C. (2017). Design and analysis of experiments (9th ed.). Wiley.

Peffers, K., Tuunanen, T., Rothenberger, M. A., & Chatterjee, S. (2007). A design science research methodology for information systems research. Journal of Management Information Systems, 24(3), 45–77. https://doi.org/10.2753/MIS0742-1222240302

Qu, C., Calheiros, R. N., & Buyya, R. (2018). Auto-scaling web applications in clouds: A taxonomy and survey. ACM Computing Surveys, 51(4), Article 73. https://doi.org/10.1145/3148149

Sargent, R. G. (2013). Verification and validation of simulation models. Journal of Simulation, 7(1), 12-24. https://doi.org/10.1057/jos.2012.20

Siachamis, G., Christodoulou, G. C., Psarakis, K., Fragkoulis, M., van Deursen, A., & Katsifodimos, A. (2024). Evaluating stream processing autoscalers. In Proceedings of the 18th ACM International Conference on Distributed and Event-Based Systems (pp. 110–122). ACM. https://doi.org/10.1145/3629104.3666036

Singh, R. P., Kumarasubramanian, B., Maheshwari, P., & Shetty, S. (2020). Auto-sizing for stream processing applications at LinkedIn. In 12th USENIX Workshop on Hot Topics in Cloud Computing. USENIX Association.

Zaharia, M., Xin, R. S., Wendell, P., Das, T., Armbrust, M., Dave, A., Meng, X., Rosen, J., Venkataraman, S., Franklin, M. J., Ghodsi, A., Gonzalez, J., Shenker, S., & Stoica, I. (2016). Apache Spark: A unified engine for big data processing. Communications of the ACM, 59(11), 56–65. https://doi.org/10.1145/2934664a

Previous IFGICT Fellow Spotlight: Vamshidhar Reddy Vemula’s Work Across Enterprise Software Engineering, Cloud Resilience and Applied ICT Research

Leave Your Comment

Working hours 

Monday – Friday from 8:30 am – 5:30 pm EST

Mon – Fri: 8AM – 5PM Saturday: 8AM – 3PM

Sunday: Closed

IFGICT World’s Largest ICT Federation​